Papers by Cassandra L. Jacobs

3 papers
Will it Unblend? (2020.findings-emnlp)

Copied to clipboard

Challenge: Blends, such as “innoventor”, are one particularly challenging class of OOV terms, as they are formed by fusing together two or more bases that relate to the intended meaning in unpredictable manners and degrees.
Approach: They propose to use a dataset of English OOV blends to quantify the difficulty of interpreting the meanings of blends by large-scale contextual language models such as BERT.
Outcome: The proposed model outperforms character-level and context-free embeddings, although their results are still far from satisfactory.
UniMorph 3.0: Universal Morphology (2020.lrec-1)

Copied to clipboard

Challenge: Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages.
Outcome: The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages.
NYTWIT: A Dataset of Novel Words in the New York Times (2020.coling-main)

Copied to clipboard

Challenge: Novel words, or Out-Of-Vocabulary words, are a pervasive problem in modern natural language processing.
Approach: They present a dataset of over 2,500 novel English words published in the New York Times . they use uncontextual and contextual predictions to predict novelty class .
Outcome: The proposed dataset includes over 2,500 novel English words published in the New York Times between November 2017 and March 2019 . baseline results show that there is room for improvement even for state-of-the-art NLP systems .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations